Papers with domain-specific representation alignment objective
ImmersiveTTS: Environment-Aware Text-to-Speech with Multimodal Diffusion Transformer and Domain-Specific Representation Alignment (2026.acl-long)
Copied to clipboard
| Challenge: | ImmersiveTTS model synthesizes intelligible speech and environmental audio from natural language descriptions. |
| Approach: | They propose an environment-aware text-to-speech model that integrates natural speech with environmental audio . the model explicitly models cross-modal interactions through a dual-stream stage . |
| Outcome: | Experimental results show that ImmersiveTTS achieves higher naturalness, intelligibility, and audio fidelity than existing approaches. |